Introduction :

ReDoSt is dedicated to the automatic detection of the DIRS1-like retrotransposons that is based both on similarity searches and the structure of the elements. 
Similarity searches are performed using alignment profiles for three charactiristic coding domains of the DIRS1-like retrotransposons : the Reverse Transcriptase (RT), the Methyltransferase (MT) and the Tyrosine Recombinase (YR). 
ReDoSt is composed of six main steps : 
(1) Identification of all the putative reverse transcriptase encoding fragments within the genome; 
(2) For each hit, extraction of the genomic hit with 5kb flanking sequence at both side, considering that all DIRS1-like elements described to date are less than 6kb in length; Within each genomic fragment retained, 
(3) tyrosine recombinase encoding domain search and 
(4) methyltransferase encoding domain search; 
(5) After obtaining the 10 kb contigs that harbor the three characteristic domains (RT, YR and MT) of DIRS1-like retrotransposons, we checked for the co-orientation and the order of these domains to discriminate other types of YR encoding retrotransposons (e.g., Ngaro and PAT elements); 
(6) At last, fragments that harbor at least two occurrences of a same domain are set aside for the copy number estimation and sequence alignments and supplementary investigations required to determine from which rearrangements (duplications or insertions) they are derived.

Sources of ReDoSt :

All the sources are available in the ReDoSt_v1.tar.gz file. Decompact the archive with the command line `tar zxvf ReDoSt_v1.tar.gz`.

The main directory (ReDoSt_v1) is organized as follows :
- the 'genomes' directory contains the genome sequences screened in fasta format (the Dictyostelium discoideum genome ('Dd') is provided as control sequences).
- the 'profiles' directory contains the alignment profiles for three coding domains of DIRS1-like retrotransposons (RT, MT and YR). 
	Three profile versions are available : 
	(i) the 'Paper' corresponds to the profiles used in Piednoel et al., 2011. 
	(ii) the 'Annotated' corresponds to the new profiles designed on the annotated element subset in Piednoel et al., 2011.
	(iii) the 'Phylogeny' corresponds to the new profiles designed on the element included in the phylogenic analysis in Piednoel et al., 2011.
- the 'results' directory, an empty directory, where the results will be written.
- the 'scripts' directory contains all the scripts needed to perform the analysis.
- the 'fasta' directory contains all the sequences used to design the 'Annotated' and 'Phylogeny' profiles.
- the 'readme' file (this file)

ReDoSt requires a stand-alone BLAST version. 'ReDoSt1.0.csh' and 'ReDoSt1.1.csh' are designed to work using BLAST and BLAST+ (2.2.24 or better), respectively. BLAST or BLAST+ have to be previously installed in the user path on the machine.

ReDoSt also requires the Python package (version 2.4 to 2.7), the NumPy package and the Biopython package.

Use of ReDoSt :

Before running ReDoSt, please ensure you that the 'DomainPath', 'ResultPath', 'ProgPath' and 'GenomePath' exist and are well defined in 'ReDoSt1.0.csh' or 'ReDoSt1.1.csh' scripts. 
Please also ensure you that the csh and python paths are well defined at the beginning of all the scripts.

Note that if the pipeline is running on a given genome that was previously already screened, you need to remove the old results contained in the 'ResultPath' for this genome.

To run the pipeline, use the command line :
ReDoSt1.x.csh name_of_the_directory_containing_the_genome_fasta_file

Example given for Dictyostelium discoideum : ReDoSt1.0.csh Dd

Note that the pipeline provided is designed to use the 'Paper' version of the alignment profiles. If you want to screen the genomes using the new 'Annotated' and 'Phylogeny' versions, you have to modify the main script ReDoSt1.x.csh using the psitblastn program instead of the rpsblast program in the 'RT search' part.

Descriptions of the files created in the ResultPath :

Several outputs are provided (example for Dictyostelium discoideum, Dd) :

- 'Dd.RT.fst' contains all the 10kb genomic sequences in fasta format where a RT has been detected.
- 'Dd.RT.YR.fst' contains all the 10kb genomic sequences in fasta format where a RT and a YR have been detected.
- 'Dd.RT.YR.MT.fst' contains all the 10kb genomic sequences in fasta format where a RT, an YR and a MT have been detected.

-'Final.Dd.RTMTYR.fst' contains all the putative nucleotide DIRS1-like retrotransposon sequences (from the beginning of the RT to the end of the YR) in fasta format.
-'Final.Dd.RTMT.fst' contains all the putative nucleotide DIRS1-like retrotransposon pol sequences (from the beginning of the RT to the end of the MT) in fasta format.
-'Final.Dd.xx.fst' contains nucleotide domain sequences of all putative DIRS1-like retrotransposons in fasta format (xx = RT, MT or YR).

-'Dd.xx.blout' : results of the similarity searches in tabular format for a coding domain (xx = RT, MT or YR).

-'annot.tmp' summarizes the similarity searche results in tabular format :
fragment_name beginningRT endRT evalueRT beginningYR endYR evalueYR beginningMT endMT evalueMT
-'annotCo.tmp' keeps only the results from 'annot.tmp' that could correspond to DIRS1-like retrotransposons in regard to the usual DIRS1-like element structure. New tabular output format :
fragment_name beginningRT endRT evalueRT beginningMT endMT evalueMT beginningYR endYR evalueYR
-'annotCoGenom.tmp', identical to 'annotCo.tmp' but provides the locations of the domain on the genome rather than on 10kb fragment.
-'annotCoGenomUniq.tmp', identical to 'annotCoGenomUniq.tmp' without the fragments where a peculiar domain occured twice or more.
-'listeDoublons.RTnumber' contains the name of the RT included in fragments where a peculiar domain occured twice or more (Format : '@RTxx@').
-'listeDoublons.tmp' contained results excluded from 'annotCoGenom.tmp' in 'annotCoGenomUniq.tmp'. Output format :
	DDD
	Fragment_name1 domain_occurrence1
	Fragment_name1 domain_occurrence2
	FFF
	DDD
	Fragment_name2 domain_occurrence1
	Fragment_name2 domain_occurrence2
	FFF


-'Dd.xx.fst.nhr', 'Dd.xx.fst.nin' and 'Dd.xx.fst.nsq' : formatted database for similarity searches (xx = RT, MT or YR).

Contact :

If you encounter some problems with ReDoSt, please contact Mathieu Piednoel (piednoel@closun.snv.jussieu.fr). Any comments or suggestions are welcome.

